iT邦幫忙

2026 iThome 鐵人賽

DAY 3
0

問題會自己露出來

昨天我們讓 Agent 成功呼叫 LLM 分析規格書。但試著丟一份 100 頁的規格書進去,你就會碰到這個:

Mistral 7B 的上下文窗口是 32k token。一份 100 頁 SRS?大約 50k+ token。LLM 會截斷,或者到中間就開始胡說八道。

簡單的解法是按字符數切:「每 5000 字符一個 chunk」。這樣能進 LLM 的上下文。但問題是,規格書不是雜文。切到一半的時候,第一個 chunk 結尾是「系統支持多用戶並行」,第二個 chunk 開頭是「...編輯、分享筆記」。LLM 看不到這兩句是同一個需求,而且它忘記了第一個 chunk 說的內容。

如果規格書的邏輯矛盾跨越了 chunk 邊界呢?LLM 看不到。

所以需要按 Markdown 層級分塊,不是按字符。

怎麼做

我們用 # 級標題作為主要章節邊界。每個 # 開始一個新 chunk,## 和 ### 跟著父 chunk。這樣:

# 第 1 章:系統概述
  ## 1.1 功能範圍
    ### 1.1.1 筆記編輯器
  ## 1.2 用戶角色
# 第 2 章:功能需求
  ## 2.1 筆記管理
    - REQ-2.1.1:系統支持多用戶

切割成:

  • Chunk 1: 第 1 章 + 所有子章節(1.1、1.1.1、1.2)
  • Chunk 2: 第 2 章 + 子章節(2.1 + 所有 REQ)

LLM 看到 REQ-2.1.1 時,知道它在「第 2 章 - 功能需求 - 筆記管理」的上下文裡。

還要防止單個 chunk 超限。我們估算 token:1 token ≈ 4 字符。簡單,夠用。

代碼

src/chunking.py:

import re
from dataclasses import dataclass
from typing import List, Optional

@dataclass
class Chunk:
    """文檔分塊"""
    content: str
    title: str
    level: int                  # 標題層級(1=#, 2=##)
    start_line: int
    end_line: int
    token_estimate: int

    def __str__(self) -> str:
        indent = "  " * (self.level - 1)
        return (
            f"{indent}{'#' * self.level} {self.title}\n"
            f"{indent}  (行 {self.start_line}-{self.end_line}, "
            f"~{self.token_estimate} tokens)"
        )

class MarkdownChunker:
    """Markdown 感知的分塊器"""

    DEFAULT_TOKEN_LIMIT = 4000
    CHARS_PER_TOKEN = 4         # 1 token ≈ 4 字符

    def __init__(self, token_limit: int = DEFAULT_TOKEN_LIMIT):
        self.token_limit = token_limit

    def _estimate_tokens(self, text: str) -> int:
        """估算 token 數"""
        return max(1, len(text) // self.CHARS_PER_TOKEN)

    def _get_heading_level(self, line: str) -> Optional[int]:
        """判斷是否是標題,返回層級或 None"""
        match = re.match(r'^(#+)\s', line)
        return len(match.group(1)) if match else None

    def chunk(self, text: str) -> List[Chunk]:
        """分塊"""
        lines = text.split('\n')
        chunks = []

        current_chunk_lines = []
        current_title = "未分類"
        current_level = 0
        current_start_line = 0

        for line_num, line in enumerate(lines):
            heading_level = self._get_heading_level(line)

            # 遇到 # 級標題,就開始新 chunk
            if heading_level == 1 and current_chunk_lines:
                chunk_text = '\n'.join(current_chunk_lines)
                chunks.append(Chunk(
                    content=chunk_text,
                    title=current_title,
                    level=current_level,
                    start_line=current_start_line,
                    end_line=line_num - 1,
                    token_estimate=self._estimate_tokens(chunk_text)
                ))
                current_chunk_lines = []
                current_start_line = line_num

            # 更新當前標題
            if heading_level:
                current_title = re.sub(r'^#+\s+', '', line).strip()
                current_level = heading_level

            current_chunk_lines.append(line)

        # 最後一個 chunk
        if current_chunk_lines:
            chunk_text = '\n'.join(current_chunk_lines)
            chunks.append(Chunk(
                content=chunk_text,
                title=current_title,
                level=current_level,
                start_line=current_start_line,
                end_line=len(lines) - 1,
                token_estimate=self._estimate_tokens(chunk_text)
            ))

        return chunks

    def validate_chunks(self, chunks: List[Chunk]) -> List[str]:
        """檢查超限"""
        warnings = []
        for i, chunk in enumerate(chunks):
            if chunk.token_estimate > self.token_limit:
                warnings.append(
                    f"⚠️  Chunk {i+1} ('{chunk.title}') "
                    f"超過限制:{chunk.token_estimate} > {self.token_limit} tokens"
                )
        return warnings

def print_chunks(chunks: List[Chunk], show_content: bool = False):
    """顯示分塊結果"""
    print(f"\n{'='*70}")
    print(f"文檔分塊結果(共 {len(chunks)} 個 chunks)")
    print(f"{'='*70}\n")

    total_tokens = 0
    for i, chunk in enumerate(chunks, 1):
        print(f"{i}. {chunk}")
        total_tokens += chunk.token_estimate

        if show_content:
            preview = chunk.content[:200].replace('\n', '\n   ')
            print(f"   預覽:{preview}...\n")
        else:
            print()

    print(f"{'='*70}")
    print(f"總計:{total_tokens} tokens")
    print(f"{'='*70}\n")

在 src/main.py 加上這個:

from src.chunking import MarkdownChunker, print_chunks

# OllamaAgent 加新方法
def chunk_srs(self, srs_text: str, token_limit: int = 4000):
    """分塊 SRS"""
    chunker = MarkdownChunker(token_limit=token_limit)
    chunks = chunker.chunk(srs_text)

    warnings = chunker.validate_chunks(chunks)
    if warnings:
        print("\n⚠️  警告:")
        for warning in warnings:
            print(f"  {warning}")

    return chunks

# 新的 CLI 命令
@app.command()
def chunk(
    filepath: str = typer.Argument(..., help="SRS 檔案路徑"),
    token_limit: int = typer.Option(
        4000,
        "--limit",
        "-l",
        help="單個 chunk 的 token 限制"),
    show_content: bool = typer.Option(
        False,
        "--content",
        "-c",
        help="是否顯示內容預覽")):
    """分塊 SRS 文檔

    使用方式:
        python -m src.main chunk tests/fixtures/srs_small.md
        python -m src.main chunk tests/fixtures/srs_small.md -c
        python -m src.main chunk tests/fixtures/srs_small.md -l 3000
    """
    try:
        agent = OllamaAgent()

        print(f"📖 讀取: {filepath}")
        srs_text = agent.read_srs(filepath)
        print(f"✓ {len(srs_text)} 字符")

        print(f"\n🔨 分塊中(限制:{token_limit} tokens/chunk)...")
        chunks = agent.chunk_srs(srs_text, token_limit=token_limit)

        print_chunks(chunks, show_content=show_content)

    except Exception as e:
        print(f"❌ 錯誤: {e}")
        raise typer.Exit(1)

測試

# 基本分塊
python -m src.main chunk tests/fixtures/srs_small.md

# 顯示預覽
python -m src.main chunk tests/fixtures/srs_small.md -c

# 降低 token 限制,看看會不會分得更多
python -m src.main chunk tests/fixtures/srs_small.md -l 2000

預期:

📖 讀取: tests/fixtures/srs_small.md
✓ 847 字符

🔨 分塊中(限制:4000 tokens/chunk)...

======================================================================
文檔分塊結果(共 4 個 chunks)
======================================================================

1. # 線上筆記系統 - 需求規格書
   (行 0-8, ~212 tokens)

2. ## 功能需求
   (行 9-16, ~280 tokens)

3. ## 性能需求
   (行 17-20, ~150 tokens)

4. ## 安全需求
   (行 21-27, ~120 tokens)

======================================================================
總計:762 tokens
======================================================================

為什麼這樣設計

按 # 級邊界,不按 ##:因為 # 是主章節,每個主章節代表一個完整的邏輯單元。如果按 ## 分,chunks 會很碎。

Token 估算用簡單公式:中文、英文都大致是 1 token ≈ 4 字符。如果想精確,可以用 tiktoken,但這裡簡單估算夠了。

明天的改進

明天(Day 4)開始,我們加入「全局約束追蹤」。現在 chunks 只是切割,它們還是獨立的。明天會讓 Agent 在分析每個 chunk 時,記住所有之前 chunks 裡提出過的約束,這樣才能檢測到「第 2 章說多用戶,第 8 章說單用戶」的矛盾。


進度

第 1 週:透明的 SRS 分析

  • Day 1:為什麼需要 Agent ✓
  • Day 2:本地 LLM 跑起來 ✓
  • Day 3:智能分塊 ← 今天
  • Day 4-5:全局約束追蹤
  • Day 6-7:MCP 工具與衝突檢測

上一篇
Day 2:30 秒內讓本地 LLM 分析你的規格書
下一篇
Day 4:為什麼 LLM 看不出矛盾
系列文
解決需求規格書矛盾:用 Claude Code × MCP 實作自律型文檔審查 Agent 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言